AITopics | recurrence and self-attention

Untangling tradeoffs between recurrence and self-attention in artificial neural networks

Neural Information Processing SystemsDec-24-2025, 19:07:23 GMT

Attention and self-attention mechanisms, are now central to state-of-the-art deep learning on sequential tasks. However, most recent progress hinges on heuristic approaches with limited understanding of attention's role in model optimization and computation, and rely on considerable memory and computational resources that scale poorly. In this work, we present a formal analysis of how self-attention affects gradient propagation in recurrent networks, and prove that it mitigates the problem of vanishing gradients when trying to capture long-term dependencies by establishing concrete bounds for gradient norms. Building on these results, we propose a relevancy screening mechanism, inspired by the cognitive process of memory consolidation, that allows for a scalable use of sparse self-attention with recurrence. While providing guarantees to avoid vanishing gradients, we use simple numerical experiments to demonstrate the tradeoffs in performance and computational resources by efficiently balancing attention and recurrence. Based on our results, we propose a concrete direction of research to improve scalability of attentive networks.

artificial neural network, recurrence and self-attention, untangling tradeoff, (6 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.58)

Add feedback

Review for NeurIPS paper: Untangling tradeoffs between recurrence and self-attention in artificial neural networks

Neural Information Processing SystemsFeb-7-2025, 08:20:08 GMT

Additional Feedback: - Line 145, how can Theorem 1 be related to the early attention mechanism [1]? As the attention weights are computed adaptively, it is unlikely that they are uniform. MANNs learn to store relevant hidden states to a fixed-size memory, which seems to have the same purpose as relevancy screening mechanism. What is the advantage of the proposed method over MANNs? How are MANNs related to the Theorem 2? - The paper neglects prior works that also aim to quantify gradient propagation in RNNs and attentive models [4,5].

neural network, recurrence and self-attention, relevancy screening mechanism, (11 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (1.00)

Add feedback

Review for NeurIPS paper: Untangling tradeoffs between recurrence and self-attention in artificial neural networks

Neural Information Processing SystemsFeb-7-2025, 08:20:00 GMT

The paper provides theoretical analysis of self-attention and vanishing gradients. Experiments are of toy problems with non-SOTA results but validate the main theoretical contributions of the paper.

artificial neural network, recurrence and self-attention, untangling tradeoff, (1 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.85)

Add feedback

Untangling tradeoffs between recurrence and self-attention in artificial neural networks

Neural Information Processing SystemsOct-11-2024, 14:26:26 GMT

Attention and self-attention mechanisms, are now central to state-of-the-art deep learning on sequential tasks. However, most recent progress hinges on heuristic approaches with limited understanding of attention's role in model optimization and computation, and rely on considerable memory and computational resources that scale poorly. In this work, we present a formal analysis of how self-attention affects gradient propagation in recurrent networks, and prove that it mitigates the problem of vanishing gradients when trying to capture long-term dependencies by establishing concrete bounds for gradient norms. Building on these results, we propose a relevancy screening mechanism, inspired by the cognitive process of memory consolidation, that allows for a scalable use of sparse self-attention with recurrence. While providing guarantees to avoid vanishing gradients, we use simple numerical experiments to demonstrate the tradeoffs in performance and computational resources by efficiently balancing attention and recurrence. Based on our results, we propose a concrete direction of research to improve scalability of attentive networks.

artificial neural network, recurrence and self-attention, untangling tradeoff, (3 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (1.00)

Add feedback

Collaborating Authors

recurrence and self-attention

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

Untangling tradeoffs between recurrence and self-attention in artificial neural networks

Review for NeurIPS paper: Untangling tradeoffs between recurrence and self-attention in artificial neural networks

Review for NeurIPS paper: Untangling tradeoffs between recurrence and self-attention in artificial neural networks

Untangling tradeoffs between recurrence and self-attention in artificial neural networks